Gene sequence signatures revealed by mining the UniGene affiliation network

نویسندگان

  • Jiexin Zhang
  • Li Zhang
  • Kevin R. Coombes
چکیده

BACKGROUND In the post-genomic era, developing tools to decode biological information from genomic sequences is important. Inspired by affiliation network theory, we investigated gene sequences of two kinds of UniGene clusters (UCs): narrowly expressed transcripts (NETs), whose expression is confined to a few tissues; and prevalently expressed transcripts (PETs) that are expressed in many tissues. RESULTS We explored the human and the mouse UniGene databases to compare NETs and PETs from different perspectives. We found that NETs were associated with smaller cluster size, shorter sequence length, a lower likelihood of having LocusLink annotations, and lower and more sporadic levels of expression. Significantly, the dinucleotide frequencies of NETs are similar to those of intergenic sequences in the genome, and they differ from those of PETs. We used these differences in dinucleotide frequencies to develop a discriminant analysis model to distinguish PETs from intergenic sequences. CONCLUSIONS Our results show that most NETs resemble intergenic sequences, casting doubts on the quality of such UniGene clusters. However, we also noted that a fraction of NETs resemble PETs in terms of dinucleotide frequencies and other features. Such NETs may have fewer quality problems. This work may be helpful in the studies of non-coding RNAs and in the validation of gene sequence databases.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Title: Gene Sequence Signatures Revealed by Mining the UniGene Affiliation Network

Background: In the post-genomic era, developing tools to decode biological information from genomic sequences is important. Inspired by affiliation network theory, we investigated gene sequences of two kinds of UniGene clusters (UCs): narrowly expressed transcripts (NETs), whose expression is confined to a few tissues; and prevalently expressed transcripts (PETs) that are expressed in many tiss...

متن کامل

Identification of differentially expressed UniGenes in developing wheat seed using Digital Differential Display

The wheat UniGene sets, derived from over one million Expressed Sequence Tags (ESTs) in the NCBI GenBank, offer a platform for identifying differentially expressed genes in wheat seeds. This report illustrates a means to efficiently utilize this public database for gene expression (transcriptome) profiling of developing wheat seed. Using a data mining tool known as Digital Differential Display ...

متن کامل

Exploring Gene Signatures in Different Molecular Subtypes of Gastric Cancer (MSS/ TP53+, MSS/TP53-): A Network-based and Machine Learning Approach

Gastric cancer (GC) is one of the leading causes of cancer mortality, worldwide. Molecular understanding of GC’s different subtypes is still dismal and it is necessary to develop new subtype-specific diagnostic and therapeutic approaches. Therefore developing comprehensive research in this area is demanding to have a deeper insight into molecular processes, underlying these subtypes. In this st...

متن کامل

UgMicroSatdb: database for mining microsatellites from unigenes

Microsatellites, also known as simple sequence repeats (SSRs) or short tandem repeats (STRs), have extensively been exploited as molecular markers for diverse applications. Recently, their role in gene regulation and genome evolution has also been discussed widely. We have developed UgMicroSatdb (Unigene MicroSatellite database), a web-based relational database of microsatellites present in uni...

متن کامل

Gene regulation network fitting of genes involved in the pathophysiology of fatty liver in the mice by promoter mining

Background and Aim: Non-Alcoholic Fatty Liver Disease (NAFLD) is the major cause of chronic liver disease in developed countries. In this study, we identified the most important transcription factors and biological mechanisms affecting the incidence of fatty liver disease using the promoter region data mining. Materials and Methods In this study, at first, the marker genes associated with this...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:
  • Bioinformatics

دوره 22 4  شماره 

صفحات  -

تاریخ انتشار 2006